heterogeneous data source
On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models
Tian, Jinchuan, Peng, Yifan, Chen, William, Choi, Kwanghee, Livescu, Karen, Watanabe, Shinji
The Open Whisper-style Speech Model (OWSM) series was introduced to achieve full transparency in building advanced speech-to-text (S2T) foundation models. To this end, OWSM models are trained on 25 public speech datasets, which are heterogeneous in multiple ways. In this study, we advance the OWSM series by introducing OWSM v3.2, which improves on prior models by investigating and addressing the impacts of this data heterogeneity. Our study begins with a detailed analysis of each dataset, from which we derive two key strategies: data filtering with proxy task to enhance data quality, and the incorporation of punctuation and true-casing using an open large language model (LLM). With all other configurations staying the same, OWSM v3.2 improves performance over the OWSM v3.1 baseline while using 15% less training data.
Development of Semantics-Based Distributed Middleware for Heterogeneous Data Integration and its Application for Drought
Drought is a complex environmental phenomenon that affects millions of people and communities all over the globe and is too elusive to be accurately predicted. This is mostly due to the scalability and variability of the web of environmental parameters that directly/indirectly causes the onset of different categories of drought. Since the dawn of man, efforts have been made to uniquely understand the natural indicators that provide signs of likely environmental events. These indicators/signs in the form of indigenous knowledge system have been used for generations. The intricate complexity of drought has, however, always been a major stumbling block for accurate drought prediction and forecasting systems. Recently, scientists in the field of agriculture and environmental monitoring have been discussing the integration of indigenous knowledge and scientific knowledge for a more accurate environmental forecasting system in order to incorporate diverse environmental information for a reliable drought forecast. Hence, in this research, the core objective is the development of a semantics-based data integration middleware that encompasses and integrates heterogeneous data models of local indigenous knowledge and sensor data towards an accurate drought forecasting system for the study areas. The local indigenous knowledge on drought gathered from the domain experts is transformed into rules to be used for performing deductive inference in conjunction with sensors data for determining the onset of drought through an automated inference generation module of the middleware. The semantic middleware incorporates, inter alia, a distributed architecture that consists of a streaming data processing engine based on Apache Kafka for real-time stream processing; a rule-based reasoning module; an ontology module for semantic representation of the knowledge bases.
The untapped potential of HPC + graph computing
In the past few years, AI has crossed the threshold from hype to reality. Today, with unstructured data growing by 23% annually in an average organization, the combination of knowledge graphs and high performance computing (HPC) is enabling organizations to exploit AI on massive datasets. Full disclosure: Before I talk about how critical graph computing HPC is going to be, I should tell you that I'm CEO of a graph computing, AI and analytics company, so I certainly have a vested interest and perspective here. But I'll also tell you that our company is one of many in this space -- DGraph, MemGraph, TigerGraph, Neo4j, Amazon Neptune, and Microsoft's CosmosDB, for example, all use some form of HPC graph computing. And there are many other graph companies and open-source graph options, including OrientDB, Titan, ArangoDB, Nebula Graph, and JanusGraph.
ROC: An Ontology for Country Responses towards COVID-19
Qundus, Jamal Al, Schäfermeier, Ralph, Karam, Naouel, Peikert, Silvio, Paschke, Adrian
The ROC ontology for country responses to COVID-19 provides a model for collecting, linking and sharing data on the COVID-19 pandemic. It follows semantic standardization (W3C standards RDF, OWL, SPARQL) for the representation of concepts and creation of vocabularies. ROC focuses on country measures and enables the integration of data from heterogeneous data sources. The proposed ontology is intended to facilitate statistical analysis to study and evaluate the effectiveness and side effects of government responses to COVID-19 in different countries. The ontology contains data collected by OxCGRT from publicly available information. This data has been compiled from information provided by ECDC for most countries, as well as from various repositories used to collect data on COVID-19.
Semantic Interoperability Middleware Architecture for Heterogeneous Environmental Data Sources
Data heterogeneity hampers the effort to integrate and infer knowledge from vast heterogeneous data sources. An application case study is described, in which the objective was to semantically represent and integrate structured data from sensor devices with unstructured data in the form of local indigenous knowledge. However, the semantic representation of these heterogeneous data sources for environmental monitoring systems is not well supported yet. To combat the incompatibility issues, a dedicated semantic middleware solution is required. In this paper, we describe and evaluate a cross-domain middleware architecture that semantically integrates and generate inference from heterogeneous data sources. These use of semantic technology for predicting and forecasting complex environmental phenomenon will increase the degree of accuracy of environmental monitoring systems.